Papers with multimodal language
Multimodal Language Analysis with Recurrent Multistage Fusion (D18-1)
Copied to clipboard
| Challenge: | Comprehending multimodal language requires modeling interactions between modalities and between them. |
| Approach: | They propose a multistage fusion network which decomposes the fusion problem into multiple stages, each focused on a subset of multimodal signals for specialized, effective fusion. |
| Outcome: | The proposed model performs state-of-the-art across three datasets relating to multimodal sentiment analysis, emotion recognition, and speaker traits recognition. |
CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French (2020.emnlp-main)
Copied to clipboard
AmirAli Bagher Zadeh, Yansheng Cao, Simon Hessner, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency
| Challenge: | Existing datasets in multimodal language are limited and disproportionately affect native speakers of other languages . authors propose a large-scale dataset for Spanish, Portuguese, German and French . |
| Approach: | They propose a large-scale multimodal language dataset for Spanish, Portuguese, German and French. |
| Outcome: | The proposed dataset is the largest of its kind with 40,000 total labelled sentences . it covers a diverse set topics and speakers and carries supervision of 20 labels including sentiment, emotions, and attributes. |
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse? (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing studies on the distribution of information in visually grounded contexts have focused on text-only inputs. |
| Approach: | They propose to use multilingual vision-and-language models to estimate surprisal . they find grounding on perception increases uniformity across typologically diverse languages . |
| Outcome: | The proposed hypothesis is tested in visual-language models over 30 languages and 13 storytelling languages . the results show grounding on perception increases uniformity across languages compared to text-only settings . |
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor (D19-1)
Copied to clipboard
Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, Mohammed (Ehsan) Hoque
| Challenge: | Humor is a unique and creative communicative behavior often displayed during social interactions. |
| Approach: | They present a dataset that allows to model multimodal language used in expressing humor using text, visual and acoustic communication. |
| Outcome: | The proposed framework opens the door to understanding multimodal language used in expressing humor. |
Integrating Multimodal Information in Large Pretrained Transformers (2020.acl-main)
Copied to clipboard
Wasifur Rahman, Md Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh, Chengfeng Mao, Louis-Philippe Morency, Ehsan Hoque
| Challenge: | Recent Transformer-based contextual word representations have shown state-of-the-art performance in multiple disciplines within NLP. |
| Approach: | They propose an attachment to BERT and XLNet that allows them to accept multimodal nonverbal data during fine-tuning. |
| Outcome: | The proposed attachment allows BERT and XLNet to accept multimodal nonverbal data during fine-tuning. |